Papers with multimodal machine translation

16 papers
Choosing What to Mask: More Informed Masking for Multimodal Machine Translation (2023.acl-srw)

Copied to clipboard

Challenge: Pre-trained language models have achieved remarkable results on several NLP tasks.
Approach: They propose three new masking strategies for cross-lingual visual pre-training that focus on learning different linguistic patterns.
Outcome: The proposed methods outperform the baseline model and achieve state-of-the-art accuracy on the Portuguese-English MMT task.
Towards Zero-Shot Multimodal Machine Translation (2025.findings-naacl)

Copied to clipboard

Challenge: Current multimodal machine translation systems rely on fully supervised data, which is costly to collect and prevents extension of MMT to language pairs with no such data.
Approach: They propose a method to bypass the need for fully supervised data to train MMT systems . they adapt a strong text-only machine translation model to a visually conditioned language model and a divergence test set to evaluate how well models use images to disambiguate translations.
Outcome: The proposed method can generalize to languages with no fully supervised training data.
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Pre-trained language models have been shown to improve performance in many natural language tasks.
Approach: They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora.
Outcome: The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation.
Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies on multimodal machine translation (MMT) have focused on the fusion and alignment of images and texts to improve MMT.
Approach: They propose an image-free inference framework that supports image-based inference via an inversion knowledge distillation scheme.
Outcome: The proposed framework is the first to rival or surpass image-must frameworks on the multimodal translation benchmark.
MSCTD: A Multimodal Sentiment Chat Translation Dataset (2022.acl-long)

Copied to clipboard

Challenge: Multimodal machine translation and textual chat translation have received considerable attention . however, little research has been devoted to multimodal machine translator in conversations .
Approach: They propose a task to generate more accurate translations with the help of dialogue history and visual context.
Outcome: The proposed task can generate more accurate translations with the help of dialogue history and visual context.
Video-Helpful Multimodal Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal machine translation datasets contain images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity.
Approach: They propose an MMT dataset that contains ambiguous subtitles and a video-helpful evaluation set.
Outcome: The proposed model performs significantly better than existing models on ambiguous subtitles dataset . it is based on a training set and video-helpful evaluation set .
Adaptive Fusion Techniques for Multimodal Data (2021.eacl-main)

Copied to clipboard

Challenge: Effective fusion of data from multiple modalities is challenging due to the heterogeneous nature of multimodal data.
Approach: They propose two adaptive fusion techniques that aim to combine multimodal data effectively.
Outcome: The proposed networks can model context from other modalities better than existing methods.
Adversarial Evaluation of Multimodal Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing evidence that visual context helps multimodal machine translation systems is unconvincing due to inconsistencies between text-similarity metrics and human judgements.
Approach: They propose an adversarial evaluation method to examine the utility of image data in multimodal machine translation.
Outcome: The proposed method shows that only one out of three publicly available systems is sensitive to this perturbation of the data.
Probing the Need for Visual Context in Multimodal Machine Translation (N19-1)

Copied to clipboard

Challenge: Current work on multimodal machine translation (MMT) suggests that the visual modality is either unnecessary or only marginally beneficial.
Approach: They propose to use the visual modality to combine visual and textual information to generate better translations by partially depriving models from source-side textual context.
Outcome: The proposed model can combine visual and textual information to generate better translations under limited textual context.
On Vision Features in Multimodal Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Recent work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is given to the quality of vision models.
Approach: They develop a selective attention model to study the patch-level contribution of an image in multimodal machine translation.
Outcome: The proposed model is able to learn translation from the visual modality on probing tasks and is compared with existing models.
BERTGen: Multi-task Generation through BERT (2021.acl-long)

Copied to clipboard

Challenge: Recent work in unsupervised and self-supervised pre-training has revolutionised the field of natural language understanding (NLU).
Approach: They propose to use multimodal and multilingual pre-trained models to extend BERT by fusing them together for language generation tasks.
Outcome: The proposed model outperforms baseline models in image captioning, machine translation and multimodal machine translation tasks and is competitive with supervised counterparts.
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking (2020.lrec-1)

Copied to clipboard

Challenge: Existing multimodal corpora lack the ability to be used in multilingual or non-English scenarios.
Approach: They extend a Flickr30k Entities image-caption dataset with Japanese translations to provide a multilingual corpus.
Outcome: The proposed dataset is the first multilingual image-caption dataset with Japanese translations.
HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for visual question answering are limited to the English language.
Approach: They present a multimodal dataset for visual question answering tasks in the Hausa language.
Outcome: The proposed dataset provides 12,044 gold standard English-Hausa parallel sentences that are semantically identical to the corresponding visual information.
Distilling Translations with Visual Awareness (P19-1)

Copied to clipboard

Challenge: Existing work on multimodal machine translation has shown that visual information is only needed in very specific cases, for example in the presence of ambiguous words where the textual context is not sufficient.
Approach: They propose a translate-and-refine approach to multimodal machine translation where images are only used by a second stage decoder to generate a good first draft translation and to improve over this draft.
Outcome: The proposed approach generates a good translation and improves over the draft by making better use of the target language textual context and making use of visual context.
VISA: An Ambiguous Subtitles Dataset for Visual Scene-aware Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Existing multimodal machine translation datasets contain images and video captions or general subtitles which rarely contain linguistic ambiguity.
Approach: They propose a dataset that consists of Japanese-English parallel sentence pairs and corresponding video clips.
Outcome: The proposed dataset is challenging for the latest MMT system and can facilitate MMT research.
Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering Pairs (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that multimodal machine translation systems exhibit decreased sensitivity to visual information when text inputs are complete.
Approach: They propose to generate parallel VQA style pairs from source text to foster more robust cross-modal interaction.
Outcome: The proposed approach generates parallel VQA style pairs from the source text, fostering more robust cross-modal interaction.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations